
昨天我們已經知道「臉部偵測與臉部網格偵測的差異」,
那你知道上面的左、右兩張圖,哪一張才是臉部網格偵測的輸出嗎?
在 App 開發時要怎麼實作這項功能呢?
今天我們就要來介紹手機螢幕背後的這位重要功臣:
「Google MediaPipe Face Landmarker」究竟這位重要功臣來自何方?有哪些厲害的同門師兄?
我們一起往下看吧!
Google AI Edge 是 Google 的裝置端 AI 開發平台,讓開發者將機器學習與生成式 AI 模型直接部署在行動裝置、網頁與嵌入式應用中,帶來更低的延遲、離線可用性與更好的資料隱私。
平台支援 TensorFlow、PyTorch、Keras、JAX 等框架(註1),
並提供 MediaPipe、LiteRT、LiteRT-LM 等工具:
註1:TensorFlow、PyTorch、Keras、JAX 都是用來設計與訓練 AI 模型的框架。這些框架產出的模型通常無法直接在手機上執行,Google AI Edge 能將它們轉換成適合在裝置端運行的格式。
MediaPipe Solutions 是 MediaPipe 開源專案中的一套 AI 工具,提供跨平台 API 與預先訓練好的模型,讓開發者能快速在 Android、iOS、網頁與 Python 應用中加入 AI 功能。
目前提供的功能涵蓋視覺、文字與音訊三大類,例如臉部偵測、手勢辨識、姿勢偵測、物件偵測、文字分類與音訊分類等。「健口動一動」App 使用的 Face Landmarker,就是其中的臉部特徵點偵測(Face landmark detection)功能。
可解決方案 Available solutions:
| 功能 | Android | Web | Python | iOS | 可客製化模型 | 功能 | Android | Web | Python | iOS | 可客製化模型 | 功能 | Android | Web | Python | iOS | 可客製化模型 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 音訊分類(Audio classification) | ● | ● | ● | ● | 影像分類(Image classification) | ● | ● | ● | ● | ● | 互動式分割(Interactive segmentation) | ● | ● | ● | ● | ||
| 臉部偵測(Face detection) | ● | ● | ● | ● | 影像嵌入(Image embedding) | ● | ● | ● | ● | 語言偵測(Language detector) | ● | ● | ● | ● | |||
| 臉部特徵點偵測(Face landmark detection) | ● | ● | ● | ● | 影像分割(Image segmentation) | ● | ● | ● | ● | 文字分類(Text classification) | ● | ● | ● | ● | ● | ||
| 手勢辨識(Gesture recognition) | ● | ● | ● | ● | ● | 物件偵測(Object detection) | ● | ● | ● | ● | ● | 文字嵌入(Text embedding) | ● | ● | ● | ● | |
| 手部特徵點偵測(Hand landmark detection) | ● | ● | ● | ● | 姿勢特徵點偵測(Pose landmark detection) | ● | ● | ● | ● | 文字校對(Text proofreading) | ● | ● | ● | ||||
| 全身特徵點偵測(Holistic landmark detection) | ● | ● | ● | ● | 大型語言模型推論(LLM Inference API) | ● | ● | 文字摘要(Text summarization) | ● | ● | ● |
註:● 表示支援;「可客製化模型」表示可使用 MediaPipe Model Maker,以自己的資料調整模型。
MediaPipe 提供線上的互動式演示平台(Instant Demos),只要打開瀏覽器,就能用視訊鏡頭或自己的圖片即時測試各項功能,並調整信心門檻、結果數量等參數。
操作示意(GIF):

Day6 App 功能清單,在延緩認知退化的開發項目有一項:「一起大冒險」,是由家人和長者共同參與的互動遊戲。
除了串 Gemini 出題,還會運用到這邊 Google MediaPipe 相關偵測功能。
(一)簡介
MediaPipe Face Landmarker 是 MediaPipe 提供的臉部特徵點偵測(Face landmark detection)功能,能在圖片、影片或即時攝影機畫面中偵測人臉,並輸出三種結果:
它由三個模型串接運作:先以 BlazeFace 模型找出人臉,再由 FaceMesh-V2 模型標出 478 個特徵點,最後由 Blendshape 模型推算表情分數。常見應用包括表情辨識、臉部濾鏡特效與虛擬頭像。
它的前身是舊版的 MediaPipe Face Mesh,2023 年升級後整合了虹膜偵測與表情分數。其中的虹膜偵測只負責找出虹膜的位置,並不具備身份辨識的功能。
(二)重要參數配置

| 官方參數(Python) | 說明 | 預設值 |
|---|---|---|
| running_mode | 執行模式:IMAGE 處理單張圖片、VIDEO 處理影片幀、LIVE_STREAM 處理攝影機等即時畫面 | IMAGE |
| num_faces | 最多偵測幾張臉。只有設為 1 時才會對結果做平滑處理 | 1 |
| min_face_detection_confidence | 「找到臉」的最低信心分數 | 0.5 |
| min_face_presence_confidence | 「確認臉還在」的最低信心分數 | 0.5 |
| min_tracking_confidence | 「持續追蹤成功」的最低信心分數 | 0.5 |
| output_face_blendshapes | 是否輸出 52 個表情分數 | False |
| output_facial_transformation_matrixes | 是否輸出臉部轉換矩陣 | False |
| result_callback | LIVE_STREAM 模式下接收結果的監聽器 | 無 |
(三)初始化:以原生端 iOS 為例
//FaceLandmarkerService.swift
final class FaceLandmarkerService: NSObject {
init?(modelPath: String) {
super.init()
let options = FaceLandmarkerOptions()
options.baseOptions.modelAssetPath = modelPath
options.runningMode = .liveStream // 執行模式
options.numFaces = 1 // 最多偵測幾張臉
options.outputFaceBlendshapes = true // 是否輸出表情分數
options.minFaceDetectionConfidence = 0.5 //「找到臉」的最低信心分數
options.minFacePresenceConfidence = 0.5 //「確認臉還在」的最低信心分數
options.minTrackingConfidence = 0.5 //「持續追蹤成功」的最低信心分數,調高會頻繁掉追蹤、重跑偵測(耗電)
options.faceLandmarkerLiveStreamDelegate = self
do {
landmarker = try FaceLandmarker(options: options)
} catch {
NSLog("[face_mesh] FaceLandmarker 建立失敗: \(error)")
return nil
}
}
}
(四)影格送入Face Landmarker
//CameraSession.swift
//相機吐影格
extension CameraSession: AVCaptureVideoDataOutputSampleBufferDelegate {
func captureOutput(_ output: AVCaptureOutput,
didOutput sampleBuffer: CMSampleBuffer,
from connection: AVCaptureConnection) {
onFrame?(sampleBuffer)
}
}
// FaceMeshPlugin.swift
public class FaceMeshPlugin: NSObject, FlutterPlugin {
private func start(detect: Bool, result: @escaping FlutterResult) {
// ...(略:載入模型、建立 FaceLandmarkerService)
// 接上相機輸出附上時間戳
// 這段 closure 在相機啟動之前就先登記好,之後每收到一格影格才執行一次
CameraSession.shared.onFrame = { [weak self] sampleBuffer in
guard let self = self else { return }
let seconds = CMTimeGetSeconds(CMSampleBufferGetPresentationTimeStamp(sampleBuffer))
self.service?.detect(sampleBuffer: sampleBuffer,
timestampMs: Int(seconds * 1000))
}
// ...(略:啟動相機)
}
}
//FaceLandmarkerService.swift
final class FaceLandmarkerService: NSObject {
private var landmarker: FaceLandmarker?
// ...(略:blendshape 與輪廓點的定義、init 參數設定)
// 送檢前先把影像寬高記錄,後續換算座標使用。
private var imageWidth = 0
private var imageHeight = 0
// 非同步送檢。
// LIVE_STREAM 模式要求 timestamp 單調遞增;
// 當 Face Landmarker 正忙於處理時,部分新影格可能被略過。
func detect(sampleBuffer: CMSampleBuffer, timestampMs: Int) {
if let pixelBuffer = CMSampleBufferGetImageBuffer(sampleBuffer) {
imageWidth = CVPixelBufferGetWidth(pixelBuffer)
imageHeight = CVPixelBufferGetHeight(pixelBuffer)
}
guard let landmarker = landmarker,
let image = try? MPImage(sampleBuffer: sampleBuffer) else { return }
try? landmarker.detectAsync(image: image, timestampInMilliseconds: timestampMs)
}
}
(五)取得輸出結果:以原生端 iOS 為例
//FaceLandmarkerService.swift
extension FaceLandmarkerService: FaceLandmarkerLiveStreamDelegate {
func faceLandmarker(_ faceLandmarker: FaceLandmarker,
didFinishDetection result: FaceLandmarkerResult?,
timestampInMilliseconds: Int,
error: Error?) {
guard let result = result,
//取得臉部特徵點(Face Landmarks)
let landmarks = result.faceLandmarks.first,
landmarks.count > 454 else {
delegate?.faceLandmarkerService(self, didProduce: nil)
return
}
var shapes: [String: Double] = [:]
//取得表情分數(Blendshapes)
if let categories = result.faceBlendshapes.first?.categories {
for category in categories {
guard let name = category.categoryName,
Self.mouthKeys.contains(name) else { continue }
shapes[name] = Double(category.score)
}
}
var contour = [Float]()
//挑出健口App需要的Landmarks,把座標值排成一維陣列
//方便過橋到Flutter能以Float32List整塊記憶體傳輸。
contour.reserveCapacity(Self.contourIndices.count * 2)
for index in Self.contourIndices {
contour.append(landmarks[index].x)
contour.append(landmarks[index].y)
}
let aspect = imageWidth > 0 ? Double(imageHeight) / Double(imageWidth) : 1.0
let frame = FaceFrame(blendshapes: shapes,
contour: contour,
imageWidth: imageWidth,
imageHeight: imageHeight,
yaw: Self.yaw(landmarks),
rollDegrees: Self.rollDegrees(landmarks, aspect),
mouthOpenRatio: Self.ratio(landmarks, 13, 14, aspect),
mouthWidthRatio: Self.ratio(landmarks, 61, 291, aspect))
delegate?.faceLandmarkerService(self, didProduce: frame)
}
}
Google AI Edge 真的是強大的門派呀!
當我點進 Google MediaPipe 的Instant Demos,真的是走進百藝峰,各種絕技看得眼花撩亂XD。尤其 Image segmentation 的 Hair Segmenter,可以很完整的分割標示我的大平頭,讓我很驚艷XDD,這也是實作、探索技術的樂趣。
話題回到今天的主題脈絡~
我們認識了Google 的裝置端 AI 開發平台「Google AI Edge」、
提供跨平台 API 與預先訓練模型的「Google MediaPipe 」、
提供的臉部特徵點偵測的「MediaPipe Face Landmarker 」。
今天也看了原生iOS端「Face Landmarker的初始化→送影格→取得輸出」的概略過程。
但是,其實Flutter有社群套件支持MediaPipe Face Landmarker,為什麼要自己橋接呢?就讓我們明天繼續看下去吧!
明天就讓我們繼續深入探索:MediaPipe Face Landmarker 吧!
感謝有緣看到這邊的你~
希望佛菩薩也祝福你:🌟平安開心 健康幸福🌟
南無觀世音菩薩🍀 南無地藏菩薩🏠 南無阿彌陀佛☀️